4.3 P-Values, Confidence Regions

1 P-Values

Suppose ϕ(X) rejects for large values of T(X). We can informally define p-value as the "under null hypothesis, probability that T(X) is as large or larger than what we observed". I.e. p(x)=PH0(T(X)≥T(x))=supθ∈Θ0Pθ(T(X)≥T(x)).

Now we give a formal definition:

P-Value

Given P,Θ0,Θ. Assume we have a test ϕα for each significance level supθ∈Θ0Eθϕα(X)≤α. (For non-randomized case, it's ϕα=1{x∈Rα})
Assume tests are monotone in α: if α1≤α2, then ϕα1(x)≤ϕα2(x). (For non-randomized case, it's Rα1⊂Rα2)
Then p-value is p(x)=sup{α|ϕα(x)<1}(=sup{α:x∉Rα}).

ϕ here measures how "extreme" an observed T(X) is.

For θ∈Θ0, Pθ(p(x)≤α)=Pθ(sup{α~|ϕα~(x)<1}≤α)≤limε→0+Pθ(ϕα+ε(x)=1)≤limε→0+(α+ε)=α.
So p-value stochastically dominates u[0,1].
If ϕα rejects for large T(X), reduces to original definition.

Note the p-value is defined based on

2 Confidence Sets

2.1 Definition

Confidence Set

C(X) is a 1−α confidence set for g(θ) if Pθ(C(X)∋g(θ)≥1−α),∀θ∈Θ.
We say C(X) covers g(θ) if C(X)∋g(θ).
Pθ(C(X)∋g(θ)) is coverage probability.
infθPθ(C(X)∋g(θ)) is confidence level.

2.2 Duality of Testing & Confidence Sets

Suppose we have a level- α test ϕ(x;a) of (2.1)H0:g(θ)=a vs H1:g(θ)≠a,∀a∈g(Θ).
We can use it to make a confidence set for g(θ):
Let C(X)={a|ϕ(x;a)<1} (all non-rejected values of θ). Then Pθ(C(X)∌g(θ))=Pθ(ϕ(x;g(θ))=1)≤α,∀θ.
Alternatively, suppose C(X) is a 1−α confidence set for g(θ). We can use C to construct a test ϕ(X) of (2.1): let ϕ(X)=1{a∉C(X)}. For θ:g(θ)=a, Eθϕ(X)=Pθ(C(X)∌g(θ))≤α. This is called inverting the test.

2.3 Confidence Interval for Median

For nonparametric model X1,⋯,Xn∼i.i.dF, (F is any c.d.f) Define g(F)=median(F)=F−1(12). Consider two-sided test H0:g(F)=μ⟺F(μ)=12 vs H1:g(F)≠μ⟺F(μ)≠12.
Denote S(X;μ)=#{Xi>μ}∼Binomial(n,1−F(μ))=H012. Reject for T(X;μ)=|S(X;μ)−n2|>cα. Then μ∈C(X)⟺|S(X;μ)−n2|≤cα⟺#{Xi>μ}∈[n2−cα,n2+cα]⟺μ∈[X(n2−cα),X(n2+cα)].

3 Confidence Intervals/Bounds

If C(X)=[C1(X),C2(X)], we say C(X) is a confidence interval (CI).

We usually get LCB/UCB by inverting a one-sided test in appropriate direction called uniformly most accurate (UMA) if test UMP. And get CI by inverting a two-sided test called UMAU if test is UMPU.